Papers with visual questions

3 papers
Finding the Evidence: Localization-aware Answer Prediction for Text Visual Question Answering (2020.coling-main)

Copied to clipboard

Challenge: Existing text VQA systems generate an answer by selecting from optical character recognition (OCR) texts or a fixed vocabulary.
Approach: They propose a localization-aware answer prediction network that generates the answer and predicts a bounding box as evidence of the generated answer.
Outcome: The proposed network outperforms existing methods on three benchmark datasets for the text VQA task by a noticeable margin.
Why Did the Chicken Cross the Road? Rephrasing and Analyzing Ambiguous Questions in VQA (2023.acl-long)

Copied to clipboard

Challenge: Visual question answering models seek to answer questions about images . ambiguity can exist at all levels of linguistic analysis, but disagreements can be difficult to detect and resolve .
Approach: They develop a question-generation model which integrates group information without supervision and uses a dataset of ambiguous examples to annotate answers.
Outcome: The proposed model can integrate answer group information without supervision and is able to fill knowledge gaps and convey requests.
Answering Cross-Dimensional Geometric Visual Questions by Multi-constraint Spatial Reasoning (2026.findings-acl)

Copied to clipboard

Challenge: Existing methods for solving complex visual questions are limited in their ability to represent in a cross-dimensional space.
Approach: They propose a method that can answer complex visual questions using cross-dimensional reasoning.
Outcome: The proposed method can answer complex visual questions in 2D to 3D space with great application value.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations